Micron Document
<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Computer audition</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Computer_audition"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Computer_audition rootpage-Computer_audition skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Computer audition</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<p><b>Computer audition</b> (<b>CA</b>) or <b>machine listening</b> is the general field of study of <a href="Algorithm" title="Algorithm">algorithms</a> and systems for audio interpretation by machines.<sup id="cite_ref-1" class="reference"><a href="#cite_note-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup> Since the notion of what it means for a machine to "hear" is very broad and somewhat vague, computer audition attempts to bring together several disciplines that originally dealt with specific problems or had a concrete application in mind. The engineer <a href="Paris_Smaragdis" title="Paris Smaragdis">Paris Smaragdis</a>, interviewed in <i><a href="MIT_Technology_Review" title="MIT Technology Review">Technology Review</a></i>, talks about these systems — "software that uses sound to locate people moving through rooms, monitor machinery for impending breakdowns, or activate traffic cameras to record accidents."<sup id="cite_ref-3" class="reference"><a href="#cite_note-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p><p>Inspired by models of <a href="Hearing_(sense)" class="mw-redirect" title="Hearing (sense)">human audition</a>, CA deals with questions of representation, <a href="Transduction_(machine_learning)" title="Transduction (machine learning)">transduction</a>, grouping, use of musical knowledge and general sound <a href="Semantics" title="Semantics">semantics</a> for the purpose of performing intelligent operations on audio and music signals by the computer. Technically this requires a combination of methods from the fields of <a href="Signal_processing" title="Signal processing">signal processing</a>, auditory modelling, music perception and <a href="Cognition" title="Cognition">cognition</a>, <a href="Pattern_recognition" title="Pattern recognition">pattern recognition</a>, and <a href="Machine_learning" title="Machine learning">machine learning</a>, as well as more traditional methods of <a href="Artificial_intelligence" title="Artificial intelligence">artificial intelligence</a> for musical knowledge representation.<sup id="cite_ref-Tanguiane1993_4-0" class="reference"><a href="#cite_note-Tanguiane1993-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup><sup id="cite_ref-Tangian1994_5-0" class="reference"><a href="#cite_note-Tangian1994-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Applications">Applications</h2></div>
<p>Like <a href="Computer_vision" title="Computer vision">computer vision</a> versus image processing, computer audition versus audio engineering deals with understanding of audio rather than processing. It also differs from problems of <a href="Speech_recognition" title="Speech recognition">speech understanding by machine</a> since it deals with general audio signals, such as natural sounds and musical recordings.
</p><p>Applications of computer audition are widely varying, and include <a href="Search_for_sounds" class="mw-redirect" title="Search for sounds">search for sounds</a>, <a href="Music_genre" title="Music genre">genre</a> recognition, acoustic monitoring, <a href="Music_transcription" class="mw-redirect" title="Music transcription">music transcription</a>, score following, <a href="Audio_texture" class="mw-redirect" title="Audio texture">audio texture</a>, <a href="Music_improvisation" class="mw-redirect" title="Music improvisation">music improvisation</a>, <a href="Speech_emotion_recognition" class="mw-redirect" title="Speech emotion recognition">emotion in audio</a> and so on.
</p>
<div class="mw-heading mw-heading2"><h2 id="Related_disciplines">Related disciplines</h2></div>
<p>Computer Audition overlaps with the following disciplines:
</p>
<ul><li><a href="Music_information_retrieval" title="Music information retrieval">Music information retrieval</a>: methods for search and analysis of similarity between music signals.</li>
<li><a href="Auditory_scene_analysis" title="Auditory scene analysis">Auditory scene analysis</a>: understanding and description of audio sources and events.</li>
<li>Computational <a href="Musicology" title="Musicology">musicology</a> and mathematical music theory: use of algorithms that employ musical knowledge for analysis of music data.</li>
<li><a href="Computer_music" title="Computer music">Computer music</a>: use of computers in creative musical applications.</li>
<li>Machine musicianship: audition driven interactive music systems.</li></ul>
<div class="mw-heading mw-heading2"><h2 id="Areas_of_study">Areas of study</h2></div>
<p>Since audio signals are interpreted by the human ear–brain system, that complex perceptual mechanism should be simulated somehow in software for "machine listening". In other words, to perform on par with humans, the computer should hear and understand audio content much as humans do. Analyzing audio accurately involves several fields: electrical engineering (spectrum analysis, filtering, and audio transforms); artificial intelligence (machine learning and sound classification);<sup id="cite_ref-6" class="reference"><a href="#cite_note-6"><span class="cite-bracket">[</span>6<span class="cite-bracket">]</span></a></sup> psychoacoustics (sound perception); cognitive sciences (neuroscience and artificial intelligence);<sup id="cite_ref-7" class="reference"><a href="#cite_note-7"><span class="cite-bracket">[</span>7<span class="cite-bracket">]</span></a></sup> acoustics (physics of sound production); and music (harmony, rhythm, and timbre). Furthermore, audio transformations such as pitch shifting, time stretching, and sound object filtering, should be perceptually and musically meaningful. For best results, these transformations require perceptual understanding of spectral models, high-level feature extraction, and sound analysis/synthesis. Finally, structuring and coding the content of an audio file (sound and metadata) could benefit from efficient compression schemes, which discard inaudible information in the sound.<sup id="cite_ref-8" class="reference"><a href="#cite_note-8"><span class="cite-bracket">[</span>8<span class="cite-bracket">]</span></a></sup> Computational models of music and sound perception and cognition can lead to a more meaningful representation, a more intuitive digital manipulation and generation of sound and music in musical human-machine interfaces.
</p><p>The study of CA could be roughly divided into the following sub-problems:
</p>
<ol><li>Representation: signal and symbolic. This aspect deals with time-frequency representations, both in terms of notes and spectral models, including pattern playback and audio texture.</li>
<li><a href="Feature_extraction" class="mw-redirect" title="Feature extraction">Feature extraction</a>: sound descriptors, segmentation, onset, <a href="Pitch_detection_algorithm" title="Pitch detection algorithm">pitch</a> and <a href="Envelope_detector" title="Envelope detector">envelope</a> detection, <a href="Pitch_class" title="Pitch class">chroma</a>, and auditory representations.</li>
<li>Musical knowledge structures: analysis of <a href="Tonality" title="Tonality">tonality</a>, <a href="Rhythm" title="Rhythm">rhythm</a>, and <a href="Harmony" title="Harmony">harmonies</a>.</li>
<li>Sound similarity: methods for comparison between sounds, sound identification, novelty detection, segmentation, and clustering.</li>
<li>Sequence modeling: matching and alignment between signals and note sequences.</li>
<li>Source separation: methods of grouping of simultaneous sounds, such as multiple pitch detection and time-frequency clustering methods.</li>
<li>Auditory cognition: modeling of emotions, anticipation and familiarity, auditory surprise, and analysis of musical structure.</li>
<li><a href="Multimodal_interaction" title="Multimodal interaction">Multi-modal</a> analysis: finding correspondences between textual, visual, and audio signals.</li></ol>
<div class="mw-heading mw-heading3"><h3 id="Representation_issues">Representation issues</h3></div>
<p>Computer audition deals with audio signals that can be represented in a variety of fashions, from direct encoding of digital audio in two or more channels to symbolically represented synthesis instructions. Audio signals are usually represented in terms of <a href="Analog_recording" title="Analog recording">analogue</a> or <a href="Digital_data" title="Digital data">digital</a> recordings. <a href="Digital_recording" title="Digital recording">Digital recordings</a> are samples of acoustic waveform or parameters of <a href="Audio_compression_(data)" class="mw-redirect" title="Audio compression (data)">audio compression</a> algorithms. One of the unique properties of musical signals is that they often combine different types of representations, such as graphical scores and sequences of performance actions that are encoded as <a href="MIDI" title="MIDI">MIDI</a> files.
</p><p>Since audio signals usually comprise multiple sound sources, then unlike speech signals that can be efficiently described in terms of specific models (such as source-filter model), it is hard to devise a <a href="Parameter" title="Parameter">parametric</a> representation for general audio. Parametric audio representations usually use <a href="Filter_bank" title="Filter bank">filter banks</a> or <a href="Sine_wave" title="Sine wave">sinusoidal</a> models to capture multiple sound parameters, sometimes increasing the representation size in order to capture internal structure in the signal. Additional types of data that are relevant for computer audition are textual descriptions of audio contents, such as annotations, reviews, and visual information in the case of audio-visual recordings.
</p>
<div class="mw-heading mw-heading3"><h3 id="Features">Features</h3></div>
<p>Description of contents of general audio signals usually requires extraction of features that capture specific aspects of the audio signal. Generally speaking, one could divide the features into signal or mathematical descriptors such as energy, description of spectral shape etc., statistical characterization such as change or novelty detection, special representations that are better adapted to the nature of musical signals or the auditory system, such as logarithmic growth of sensitivity (<a href="Bandwidth_(signal_processing)" title="Bandwidth (signal processing)">bandwidth</a>) in frequency or <a href="Octave" title="Octave">octave</a> invariance (chroma).
</p><p>Since parametric models in audio usually require very many parameters, the features are used to summarize properties of multiple parameters in a more compact or salient representation.
</p>
<div class="mw-heading mw-heading3"><h3 id="Musical_knowledge">Musical knowledge</h3></div>
<p>Finding specific musical structures is possible by using musical knowledge as well as supervised and unsupervised machine learning methods. Examples of this include detection of tonality according to distribution of frequencies that correspond to patterns of occurrence of notes in musical scales, distribution of note onset times for detection of beat structure, distribution of energies in different frequencies to detect musical chords and so on.
</p>
<div class="mw-heading mw-heading3"><h3 id="Sound_similarity_and_sequence_modeling">Sound similarity and sequence modeling</h3></div>
<p>Comparison of sounds can be done by comparison of features with or without reference to time. In some cases an overall similarity can be assessed by close values of features between two sounds. In other cases when temporal structure is important, methods of dynamic time warping need to be applied to "correct" for different temporal scales of acoustic events. Finding repetitions and similar sub-sequences of sonic events is important for tasks such as texture synthesis and <a href="Machine_improvisation" class="mw-redirect" title="Machine improvisation">machine improvisation</a>.
</p>
<div class="mw-heading mw-heading3"><h3 id="Source_separation">Source separation</h3></div>
<p>Since one of the basic characteristics of general audio is that it comprises multiple simultaneously sounding sources, such as multiple musical instruments, people talking, machine noises or animal vocalization, the ability to identify and separate individual sources is very desirable. Unfortunately, there are no methods that can solve this problem in a <a href="https://en.wiktionary.org/wiki/robust" class="extiw external" title="wikt:robust">robust</a> fashion. Existing methods of source separation rely sometimes on correlation between different audio channels in <a href="Multi-channel_recording" class="mw-redirect" title="Multi-channel recording">multi-channel recordings</a>. The ability to separate sources from stereo signals requires different techniques than those usually applied in communications where multiple sensors are available. Other source separation methods rely on training or clustering of features in mono recording, such as tracking harmonically related partials for multiple pitch detection. Some methods, before explicit recognition, rely on revealing structures in data without knowing the structures (like recognizing objects in abstract pictures without attributing them meaningful labels) by finding the least complex data representations, for instance describing audio scenes as generated by a few tone patterns and their trajectories (polyphonic voices) and acoustical contours drawn by a tone (chords).<sup id="cite_ref-Tanguiane1995_9-0" class="reference"><a href="#cite_note-Tanguiane1995-9"><span class="cite-bracket">[</span>9<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading3"><h3 id="Auditory_cognition">Auditory cognition</h3></div>
<p>Listening to music and general audio is commonly not a task directed activity. People enjoy music for various poorly understood reasons, which are commonly referred to the <a href="Music_and_emotion" title="Music and emotion">emotional effect of music</a> due to creation of expectations and their realization or violation. Animals attend to signs of danger in sounds, which could be either specific or general notions of surprising and unexpected change. Generally, this creates a situation where computer audition can not rely solely on detection of specific features or sound properties and has to come up with general methods of adapting to changing auditory environment and monitoring its structure. This consists of analysis of larger repetition and <a href="Self-similarity" title="Self-similarity">self-similarity</a> structures in audio to detect innovation, as well as ability to predict local feature dynamics.
</p>
<div class="mw-heading mw-heading3"><h3 id="Multi-modal_analysis">Multi-modal analysis</h3></div>
<p>Among the available data for describing music, there are textual representations, such as liner notes, reviews and criticisms that describe the audio contents in words. In other cases human reactions such as emotional judgements or psycho-physiological measurements might provide an insight into the contents and structure of audio. Computer Audition tries to find relation between these different representations in order to provide this additional understanding of the audio contents.
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<ul><li><a href="3D_sound_localization" title="3D sound localization">3D sound localization</a></li>
<li><a href="Audio_signal_processing" title="Audio signal processing">Audio signal processing</a></li>
<li><a href="List_of_emerging_technologies" title="List of emerging technologies">List of emerging technologies</a></li>
<li><a href="Medical_intelligence_and_language_engineering_lab" title="Medical intelligence and language engineering lab">Medical intelligence and language engineering lab</a></li>
<li><a href="Music_and_artificial_intelligence" title="Music and artificial intelligence">Music and artificial intelligence</a></li>
<li><a href="Sound_recognition" title="Sound recognition">Sound recognition</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="External_links">External links</h2></div>
<ul><li><a rel="nofollow" class="external text" href="http://eceweb.ucsd.edu/~gert/calab/">UCSD Computer Audition Lab </a></li>
<li><a rel="nofollow" class="external text" href="http://www.cs.uvic.ca/~gtzan/work/caudition.html">George Tzanetakis' Computer Audition Resources</a></li>
<li><a rel="nofollow" class="external text" href="http://music.ucsd.edu/~sdubnov/ComputerAudition.htm">Shlomo Dubnov's Tutorial on Computer Audition</a></li>
<li><a rel="nofollow" class="external text" href="http://www.ee.iisc.ernet.in/">Department of Electrical Engineering, IIT (Bangalore)</a></li>
<li><a rel="nofollow" class="external text" href="https://www.smc.aau.dk/">Sound and Music Computing, Aalborg University Copenhagen, Denmark</a></li></ul>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */


.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}


/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap"><ol class="references">
<li id="cite_note-1"><span class="mw-cite-backlink"><b><a href="#cite_ref-1">^</a></b></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */


.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}


/* end https://en.wikipedia.org/ */
</style><cite class="citation book cs1"><a rel="nofollow" class="external text" href="http://www.igi-global.com/book/machine-audition-principles-algorithms-systems/40288"><i>Machine Audition: Principles, Algorithms and Systems</i></a>. IGI Global. 2011. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>9781615209194</bdi>.</cite></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><cite class="citation web cs1"><a rel="nofollow" class="external text" href="http://epubs.surrey.ac.uk/596085/1/Wang_Preface_MA_2010.pdf">"Machine Audition: Principles, Algorithms and Systems"</a> <span class="cs1-format">(PDF)</span>.</cite></span>
</li>
<li id="cite_note-3"><span class="mw-cite-backlink"><b><a href="#cite_ref-3">^</a></b></span> <span class="reference-text"><a rel="nofollow" class="external text" href="http://www.technologyreview.com/blog/VideoPosts.aspx?id=17438">Paris Smaragdis taught computers how to play more life-like music</a></span>
</li>
<li id="cite_note-Tanguiane1993-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-Tanguiane1993_4-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFTanguiane_(Tangian)1993" class="citation book cs1">Tanguiane (Tangian), Andranick (1993). <i>Artificial Perception and Music Recognition</i>. Lecture Notes in Artificial Intelligence. Vol.&nbsp;746. Berlin-Heidelberg: Springer. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a>&nbsp;<bdi>978-3-540-57394-4</bdi>.</cite></span>
</li>
<li id="cite_note-Tangian1994-5"><span class="mw-cite-backlink"><b><a href="#cite_ref-Tangian1994_5-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFTanguiane_(Tanguiane)1994" class="citation journal cs1">Tanguiane (Tanguiane), Andranick (1994). "A principle of correlativity of perception and its application to music recognition". <i>Music Perception</i>. <b>11</b> (4): <span class="nowrap">465–</span>502. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.2307%2F40285634">10.2307/40285634</a>. <a href="JSTOR_(identifier)" class="mw-redirect" title="JSTOR (identifier)">JSTOR</a>&nbsp;<a rel="nofollow" class="external text" href="https://www.jstor.org/stable/40285634">40285634</a>.</cite></span>
</li>
<li id="cite_note-6"><span class="mw-cite-backlink"><b><a href="#cite_ref-6">^</a></b></span> <span class="reference-text"><cite id="CITEREFKellyCaulfield2015" class="citation journal cs1">Kelly, Daniel; Caulfield, Brian (Feb 2015). <a rel="nofollow" class="external text" href="https://pure.ulster.ac.uk/en/publications/pervasive-sound-sensing-a-weakly-supervised-training-approach-3">"Pervasive Sound Sensing: A Weakly Supervised Training Approach"</a>. <i>IEEE Transactions on Cybernetics</i>. <b>46</b> (1): <span class="nowrap">123–</span>135. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1109%2FTCYB.2015.2396291">10.1109/TCYB.2015.2396291</a>. <a href="Hdl_(identifier)" class="mw-redirect" title="Hdl (identifier)">hdl</a>:<span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://hdl.handle.net/10197%2F6853">10197/6853</a></span>. <a href="PMID_(identifier)" class="mw-redirect" title="PMID (identifier)">PMID</a>&nbsp;<a rel="nofollow" class="external text" href="https://pubmed.ncbi.nlm.nih.gov/25675471">25675471</a>. <a href="S2CID_(identifier)" class="mw-redirect" title="S2CID (identifier)">S2CID</a>&nbsp;<a rel="nofollow" class="external text" href="https://api.semanticscholar.org/CorpusID:16042016">16042016</a>.</cite></span>
</li>
<li id="cite_note-7"><span class="mw-cite-backlink"><b><a href="#cite_ref-7">^</a></b></span> <span class="reference-text">Hendrik Purwins, Perfecto Herrera, Maarten Grachten, Amaury Hazan, Ricard Marxer, and Xavier Serra. Computational models of music perception and cognition I: The perceptual and cognitive processing chain. Physics of Life Reviews, vol. 5, no. 3, pp. 151-168, 2008. <a rel="nofollow" class="external autonumber" href="http://www.mtg.upf.edu/node/938">[1]</a></span>
</li>
<li id="cite_note-8"><span class="mw-cite-backlink"><b><a href="#cite_ref-8">^</a></b></span> <span class="reference-text"><a rel="nofollow" class="external text" href="http://web.media.mit.edu/~tristan/Classes/MAS.945/technical.html">Machine Listening Course Webpage at MIT</a></span>
</li>
<li id="cite_note-Tanguiane1995-9"><span class="mw-cite-backlink"><b><a href="#cite_ref-Tanguiane1995_9-0">^</a></b></span> <span class="reference-text"><cite id="CITEREFTanguiane_(Tangian)1995" class="citation journal cs1">Tanguiane (Tangian), Andranick (1995). "Towards axiomatization of music perception". <i>Journal of New Music Research</i>. <b>24</b> (3): <span class="nowrap">247–</span>281. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1080%2F09298219508570685">10.1080/09298219508570685</a>.</cite></span>
</li>
</ol></div></div>
<div class="navbox-styles"><style data-mw-deduplicate="TemplateStyles:r1129693374">
/* start https://en.wikipedia.org/ */


.mw-parser-output .hlist dl,.mw-parser-output .hlist ol,.mw-parser-output .hlist ul{margin:0;padding:0}.mw-parser-output .hlist dd,.mw-parser-output .hlist dt,.mw-parser-output .hlist li{margin:0;display:inline}.mw-parser-output .hlist.inline,.mw-parser-output .hlist.inline dl,.mw-parser-output .hlist.inline ol,.mw-parser-output .hlist.inline ul,.mw-parser-output .hlist dl dl,.mw-parser-output .hlist dl ol,.mw-parser-output .hlist dl ul,.mw-parser-output .hlist ol dl,.mw-parser-output .hlist ol ol,.mw-parser-output .hlist ol ul,.mw-parser-output .hlist ul dl,.mw-parser-output .hlist ul ol,.mw-parser-output .hlist ul ul{display:inline}.mw-parser-output .hlist .mw-empty-li{display:none}.mw-parser-output .hlist dt::after{content:": "}.mw-parser-output .hlist dd::after,.mw-parser-output .hlist li::after{content:" · ";font-weight:bold}.mw-parser-output .hlist dd:last-child::after,.mw-parser-output .hlist dt:last-child::after,.mw-parser-output .hlist li:last-child::after{content:none}.mw-parser-output .hlist dd dd:first-child::before,.mw-parser-output .hlist dd dt:first-child::before,.mw-parser-output .hlist dd li:first-child::before,.mw-parser-output .hlist dt dd:first-child::before,.mw-parser-output .hlist dt dt:first-child::before,.mw-parser-output .hlist dt li:first-child::before,.mw-parser-output .hlist li dd:first-child::before,.mw-parser-output .hlist li dt:first-child::before,.mw-parser-output .hlist li li:first-child::before{content:" (";font-weight:normal}.mw-parser-output .hlist dd dd:last-child::after,.mw-parser-output .hlist dd dt:last-child::after,.mw-parser-output .hlist dd li:last-child::after,.mw-parser-output .hlist dt dd:last-child::after,.mw-parser-output .hlist dt dt:last-child::after,.mw-parser-output .hlist dt li:last-child::after,.mw-parser-output .hlist li dd:last-child::after,.mw-parser-output .hlist li dt:last-child::after,.mw-parser-output .hlist li li:last-child::after{content:")";font-weight:normal}.mw-parser-output .hlist ol{counter-reset:listitem}.mw-parser-output .hlist ol>li{counter-increment:listitem}.mw-parser-output .hlist ol>li::before{content:" "counter(listitem)"\a0 "}.mw-parser-output .hlist dd ol>li:first-child::before,.mw-parser-output .hlist dt ol>li:first-child::before,.mw-parser-output .hlist li ol>li:first-child::before{content:" ("counter(listitem)"\a0 "}


/* end https://en.wikipedia.org/ */
</style><style data-mw-deduplicate="TemplateStyles:r1236075235">
/* start https://en.wikipedia.org/ */


.mw-parser-output .navbox{box-sizing:border-box;border:1px solid #a2a9b1;width:100%;clear:both;font-size:88%;text-align:center;padding:1px;margin:1em auto 0}.mw-parser-output .navbox .navbox{margin-top:0}.mw-parser-output .navbox+.navbox,.mw-parser-output .navbox+.navbox-styles+.navbox{margin-top:-1px}.mw-parser-output .navbox-inner,.mw-parser-output .navbox-subgroup{width:100%}.mw-parser-output .navbox-group,.mw-parser-output .navbox-title,.mw-parser-output .navbox-abovebelow{padding:0.25em 1em;line-height:1.5em;text-align:center}.mw-parser-output .navbox-group{white-space:nowrap;text-align:right}.mw-parser-output .navbox,.mw-parser-output .navbox-subgroup{background-color:#fdfdfd}.mw-parser-output .navbox-list{line-height:1.5em;border-color:#fdfdfd}.mw-parser-output .navbox-list-with-group{text-align:left;border-left-width:2px;border-left-style:solid}.mw-parser-output tr+tr>.navbox-abovebelow,.mw-parser-output tr+tr>.navbox-group,.mw-parser-output tr+tr>.navbox-image,.mw-parser-output tr+tr>.navbox-list{border-top:2px solid #fdfdfd}.mw-parser-output .navbox-title{background-color:#ccf}.mw-parser-output .navbox-abovebelow,.mw-parser-output .navbox-group,.mw-parser-output .navbox-subgroup .navbox-title{background-color:#ddf}.mw-parser-output .navbox-subgroup .navbox-group,.mw-parser-output .navbox-subgroup .navbox-abovebelow{background-color:#e6e6ff}.mw-parser-output .navbox-even{background-color:#f7f7f7}.mw-parser-output .navbox-odd{background-color:transparent}.mw-parser-output .navbox .hlist td dl,.mw-parser-output .navbox .hlist td ol,.mw-parser-output .navbox .hlist td ul,.mw-parser-output .navbox td.hlist dl,.mw-parser-output .navbox td.hlist ol,.mw-parser-output .navbox td.hlist ul{padding:0.125em 0}.mw-parser-output .navbox .navbar{display:block;font-size:100%}.mw-parser-output .navbox-title .navbar{float:left;text-align:left;margin-right:0.5em}body.skin--responsive .mw-parser-output .navbox-image img{max-width:none!important}@media print{body.ns-0 .mw-parser-output .navbox{display:none!important}}


/* end https://en.wikipedia.org/ */
</style></div><div role="navigation" class="navbox" aria-labelledby="Computer_audition21" style="padding:3px"><table class="nowraplinks mw-collapsible autocollapse navbox-inner" style="border-spacing:0;background:transparent;color:inherit"><tbody><tr><th scope="col" class="navbox-title" colspan="2"><style data-mw-deduplicate="TemplateStyles:r1239400231">
/* start https://en.wikipedia.org/ */


.mw-parser-output .navbar{display:inline;font-size:88%;font-weight:normal}.mw-parser-output .navbar-collapse{float:left;text-align:left}.mw-parser-output .navbar-boxtext{word-spacing:0}.mw-parser-output .navbar ul{display:inline-block;white-space:nowrap;line-height:inherit}.mw-parser-output .navbar-brackets::before{margin-right:-0.125em;content:"[ "}.mw-parser-output .navbar-brackets::after{margin-left:-0.125em;content:" ]"}.mw-parser-output .navbar li{word-spacing:-0.125em}.mw-parser-output .navbar a>span,.mw-parser-output .navbar a>abbr{text-decoration:inherit}.mw-parser-output .navbar-mini abbr{font-variant:small-caps;border-bottom:none;text-decoration:none;cursor:inherit}.mw-parser-output .navbar-ct-full{font-size:114%;margin:0 7em}.mw-parser-output .navbar-ct-mini{font-size:114%;margin:0 4em}html.skin-theme-clientpref-night .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}@media(prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .navbar li a abbr{color:var(--color-base)!important}}@media print{.mw-parser-output .navbar{display:none!important}}


/* end https://en.wikipedia.org/ */
</style><div id="Computer_audition21" style="font-size:114%;margin:0 4em"></div></th></tr><tr><td colspan="2" class="navbox-list navbox-odd hlist" style="width:100%;padding:0"><div style="padding:0 0.25em">
<ul><li><a href="Acoustic_fingerprint" title="Acoustic fingerprint">Acoustic fingerprint</a></li>
<li><a href="Audio_mining" title="Audio mining">Audio mining</a></li>
<li><a href="Computational_auditory_scene_analysis" title="Computational auditory scene analysis">Computational auditory scene analysis</a></li>
<li><a href="Music_information_retrieval" title="Music information retrieval">Music information retrieval</a></li>
<li><a href="Semantic_audio" title="Semantic audio">Semantic audio</a></li>
<li><a href="Speech_processing" title="Speech processing">Speech processing</a>
<ul><li><a href="Speech_analytics" title="Speech analytics">Speech analytics</a></li>
<li><a href="Speaker_recognition" title="Speaker recognition">Speaker recognition</a></li>
<li><a href="Speech_recognition" title="Speech recognition">Speech recognition</a></li></ul></li>
<li><a href="Sound_recognition" title="Sound recognition">Sound recognition</a></li>
<li><a href="3D_sound_localization" title="3D sound localization">3D sound localization</a></li>
<li><a href="3D_sound_reconstruction" title="3D sound reconstruction">3D sound reconstruction</a></li></ul>
</div></td></tr></tbody></table></div></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2024-03-07" href="https://en.wikipedia.org/wiki/?title=Computer_audition&amp;oldid=1212329898">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>

</body></html>